Back

Journal of Neuroscience Methods

Elsevier BV

Preprints posted in the last 7 days, ranked by how well they match Journal of Neuroscience Methods's content profile, based on 122 papers previously published here. The average preprint has a 0.09% match score for this journal, so anything above that is already an above-average fit.

1
Patient-Specific EEG Baseline Establishment Using the E-norms Method for Pediatric Seizure Detection Without Labeled Training Data

Jabre, J. F.

2026-07-16 neurology 10.64898/2026.07.13.26357876 medRxiv
Top 0.4%
4.8%
Show abstract

The aim of this work is to validate patient-specific EEG baseline establishment using the e-norms method as a screening and retrospective-review tool for seizure detection in pediatric epilepsy. The method was applied to 247 seizure-free EEG recordings (263.92 hours) from 10 patients in the CHB-MIT Scalp EEG Database (ages 3-18). A composite stability metric combining first-derivative dynamics, spectral entropy, variance, and line length was computed per 2-second epoch across 23 channels. Patient-specific detection thresholds were derived from each patient's seizure-free baseline using a weighted statistical procedure. Performance was validated against 72 expert-annotated seizures (2,705 epochs) across 62 seizure files, with durations spanning 6 to 264 seconds (44-fold range). The results show that detection achieved 94.4% event-level sensitivity (68 of 72 seizures; 95% CI 86.6-97.8%) and 81.5% epoch-level sensitivity (2,204 of 2,705 epochs; 95% CI 80.0-82.9%). Eight of ten patients achieved 100% event-level sensitivity with epoch-level sensitivity ranging from 58.7% to 100.0%. Two patients showed partial event-level failures (CHB-15: 17 of 20; CHB-18: 5 of 6), with the four missed events attributable to two characterizable failure modes. Patient-specific thresholds ranged from 4.06 to 4.81 (mean 4.51 +/- 0.25); threshold variation did not correlate reliably with age or sex, confirming that no universal threshold could achieve comparable performance. Detection margins ranged from 0.88 to 1.24 times. Patient-specific e-norms achieves 94.4% event-level sensitivity for pediatric EEG seizure detection without requiring labeled seizure training data, exceeding published human expert inter-rater agreement (50-76%) and recent automated approaches in adult cohorts using behind-the-ear EEG and wearable ECG. Two characterizable failure modes account for the four missed events and inform appropriate clinical use. As a high-sensitivity screening tool complementary to real-time alarm systems, the method is ready for adult validation, prospective deployment, and head-to-head benchmarking.

2
Initial Technical and Clinical Validation of Mobile Pupillometry with Virtual Reality: A Digital Biomarker for Screening Cognitive Function and Impairment

Brendler, A.; Fietz, J.; Bauer, A.; Pfahl, D.; Higgins, S.; Vidovic, E.; Brueckl, T.; BeCOME Working Group, ; Memory Clinic Working Group, ; Hupe, K.; Knop, M.; Spoormaker, V. I.

2026-07-17 neurology 10.64898/2026.07.15.26358187 medRxiv
Top 2%
0.8%
Show abstract

Cognitive impairment is a prevalent symptom extending from physiological ageing to disease. It commonly manifests itself in initial memory problems, progressing and co-occurring in more severe conditions such as Mild Cognitive Impairment, Alzheimer's Disease and Major Depressive Disorder. However, current non-invasive screening assessments either lack biological information or are invasive and restricted to specialized centers with complex and cost-intensive set-ups. Here, we conducted an initial validation of mobile pupillometry with Virtual Reality (VR) under experimental conditions as a digital biomarker for cognitive impairment by testing required biomarker-specific properties. For this purpose, we first assessed its construct validity by testing healthy participants (n=43) on an n-back task in VR while pupil size was measured. Mixed effects models revealed that similar to lab-based eye-tracking systems, pupil size increased in a sensible and distinguishable fashion as a function of working memory load. Second, to test the signal's reliability, the same participants were tested on the identical set-up two to three months after their first visit. We observed that the pupil response profile was highly stable over this period. Third, for its clinical validity, we examined patients (n=89) from three different cohorts with varying degrees of cognitive impairment and compared them to healthy control participants (n=81). Mixed-effects models indicated that pupil size was reduced as a function of cognitive impairment levels at higher cognitive load and that this effect was stronger pronounced with increasing age. In conclusion, we provide initial evidence for mobile pupillometry being a sensitive, reliable and clinically valid digital biomarker for cognitive functioning and impairment, which offers desirable properties due to its quick, automatized and location-independent set-up. Keywords: digital biomarker, mobile pupillometry, Virtual Reality, cognition, , Major Depressive Disorder, Mild Cognitive Impairment, Alzheimer's Disease

3
Validity and Reliability of the Novel Indonesian Instrument for Aphasia Diagnosis (IDEA)

Prawiroharjo, P.; Fakhri, A.; Gabrielle, A.; Martalia, V.; Rahmayani, S. A.; Wijaya, V. G.

2026-07-19 neurology 10.64898/2026.07.17.26358303 medRxiv
Top 4%
0.2%
Show abstract

Aphasia diagnosis in Indonesia remains challenging due to limited culturally and linguistically appropriate instruments. Widely used tools such as the Boston Diagnostic Aphasia Examination (BDAE) and Western Aphasia Battery (WAB) are not adapted to the Indonesian context, while Tes Afasia untuk Diagnosis, Informasi, dan Rehabilitasi (TADIR) provides screening but lacks diagnostic accuracy. To address this gap, we developed the Instrumen Diagnosis dan Evaluasi Afasia (IDEA) for native Indonesian speakers and evaluated its validity, reliability, and normative cutoff values in cognitively healthy Indonesian adults. Eighty-three cognitively normal adults (screened using MoCA-Ina) with no history of neurological disease were assessed using IDEA, which evaluates six language domains. Items were adapted from existing tools and reviewed by experts. Content validity, internal consistency (Cronbachs alpha), and construct validity (Exploratory Factor Analysis) were analyzed using SPSS v25. A total of 83 participants were included (median age = 55.81 years, 54% secondary education). IDEA demonstrated good feasibility, with an average completion time of 45-60 minutes depending on participant engagement. Content validity was established by unanimous expert consensus. Construct validity showed meritorious sampling adequacy (KMO = .872) and significant sphericity (Bartletts test {chi}^2 (15) = 278.523, p<.001), supporting factor analysis. Internal consistency showed good reliability across six domains (Cronbachs = 0.896). IDEA is a valid and reliable tool for assessing aphasia in Indonesian natives. It is a culturally appropriate assessment tool which offers structured, domain-based evaluation and supports differential diagnosis of both classical and progressive aphasia syndromes. Keywords: Aphasia, Language Assessment, Indonesian, IDEA, Validity

4
A Culturally Embedded Augmented Reality Task as a Neurocognitive Biomarker of Executive Function in Schizophrenia

Chatthong, W.; Rueankam, M.; Khemthong, S.

2026-07-16 psychiatry and clinical psychology 10.64898/2026.07.14.26358053 medRxiv
Top 4%
0.2%
Show abstract

Executive function (EF) deficits are central features of schizophrenia and strongly influence long-term functional outcomes. Conventional cognitive assessments often lack ecological validity and cultural relevance. This study introduces the Luk Chup Augmented Reality (LCAR) tool a video guided, clay modeling task delivered through wearable AR that integrates culturally familiar activity with realtime neurophysiological monitoring. Thirty individuals diagnosed with schizophrenia (mean age = 38.9, SD. = 7.15 years) completed a series of modeling and memory tasks using LCAR while undergoing quantitative EEG (QEEG). Task duration and theta/beta power were analyzed across procedural and color shape memory phases. Memory phases took significantly longer to complete and were associated with decreased lateral prefrontal theta and increased frontal midline theta activity (Fz, Cz), indicating higher EF demand. A repeated-measures ANOVA revealed significant condition, site, and interaction effects on theta power. The LCAR tool shows promise as a culturally grounded, dual-mode assessment of EF in schizophrenia. It offers a novel integration of performance-based and neurophysiological metrics that may inform future interventions in psychiatric rehabilitation.

5
Stereoelectroencephalography accuracy in a series of over 3000 trajectories

Thurairajah, A.; Gilmore, G.; Persad, A. R.; Youshani, A. S.; Taha, A.; Abbass, M.; Santyr, B.; Al-Orabi, K. M.; Burneo, J. G.; Pellegrino, G.; Suller-Marti, A.; Western Epilepsy Research Group, ; Parrent, A. G.; MacDougall, K. W.; Steven, D. A.; Lau, J. C.

2026-07-16 surgery 10.64898/2026.07.14.26358071 medRxiv
Top 5%
0.2%
Show abstract

Background and Objectives: Stereoelectroencephalography (SEEG) involves the implantation of intracerebral electrodes to investigate drug-resistant epilepsy. SEEG requires millimetric accuracy to ensure safety and optimal mapping. Although studies have evaluated SEEG accuracy, there is substantial variability in reporting. Here we report on implantation accuracy in a large series using the most common accuracy metrics described in the literature and perform a detailed analysis of contributing factors. Methods: SEEG implantations between 2013 and 2025 were included. Application accuracy was computed for each implanted electrode. Specifically, Euclidean, radial, depth, and angle error were calculated at both target and entry points. Correlative and multivariable analyses were conducted between each variable and error metric. Trajectories were also grouped by atlas-derived lobar target. Results: No metrics met assumptions of normality and thus we report accuracy using median with interquartile range (IQR). In a series of 3176 trajectories, median Euclidean target and entry errors were lower for robot-assisted electrodes (n=2858) at 2.19 (IQR: 1.54-2.98) mm and 1.38 (IQR: 0.89-2.01) mm respectively, compared to frame-based (n=318, p<.001) at 2.76 (IQR:1.79-3.76) mm and 2.21 (IQR: 1.42-3.32) mm. Correlation and multivariable regression analysis showed target error was positively correlated with implantation angle, scalp thickness, skull thickness, and trajectory length. Target error was also higher in obese patients. On lobar analysis, parietal lobe trajectories were the most accurate and frontal lobe trajectories were the least accurate. On temporal lobe trajectory analysis, posterior hippocampus trajectories were the most accurate and temporal pole trajectories were the least accurate. Presence of mesial temporal sclerosis also impacted accuracy. Conclusions: We present a detailed description of SEEG implantation accuracy, demonstrating the superior accuracy and speed of robot-assisted to frame-based methods. Furthermore, we analyzed how accuracy varies with specific factors from a global to trajectory level, which can be accounted for when planning SEEG implantations.

6
Statistical Inference and Power Analysis for Comparative F1 and Fβ Scores under Correlated Classifier Pairs

Hsu, C.-Y.; Liu, Q.; Shyr, Y.

2026-07-17 dermatology 10.64898/2026.07.15.26358166 medRxiv
Top 5%
0.2%
Show abstract

As machine learning and artificial intelligence systems are increasingly used in healthcare, rigorous evaluation of their classification performance has become critical. The F1 and F{beta} scores are widely adopted metrics for assessing performance in imbalanced biomedical data. Recently, we introduced psF1, a unified statistical framework for inference and study design for single and comparative F1 and F{beta} scores under the assumption of independent classifiers. In practice, however, benchmarking two classifiers on the same dataset creates a correlated paired setting. Ignoring this intrinsic dependency leads to overestimation of the standard error and a substantial loss of statistical power. To address this, we develop psF1pair, an advanced framework for statistical inference and power analysis that explicitly accounts for correlations between classifier pairs. Extensive simulation studies demonstrate the performance of psF1pair, and its utility is further illustrated through application to a real-world imaging classification system. As expected, higher correlation between classifiers yields narrower confidence intervals and enhanced statistical power. A freely available R package is provided to facilitate implementation, supporting accurate evaluation and study design for predictive and classification models in biomedical research.

7
Large Language Model - Enhanced Decision Tree Framework for Identifying Multiple Sclerosis Diagnoses from Clinical Documentation

Venkatesh, S.; DelSignore, M.; Wu, X.; Morris, M.; Kerr, W. T.; Visweswaran, S.; Wang, Y.; Xia, Z.

2026-07-17 neurology 10.64898/2026.07.14.26357416 medRxiv
Top 5%
0.1%
Show abstract

Background. Early diagnosis and intervention are crucial in multiple sclerosis (MS), yet diagnostic delays are common. Large language models (LLMs) such as generative pre-trained transformers (GPTs) may help streamline diagnostic workflows by extracting MS diagnostic signals from clinical notes. Objective. To derive MS diagnosis status from the first neurology note using a computable algorithm based on the 2017 McDonald criteria and applying GPT-4 for node-level reasoning within a structured decision framework. Methods. We analyzed first neurology notes from 125 randomly selected patients (including those with MS, related disorders, and controls) enrolled in a clinic cohort between 2017 and 2023. We included the clinical history and diagnostic testing sections but redacted the assessment and plan. We converted the 2017 McDonald criteria into a decision tree and provided expert-curated clinical knowledge to guide GPT-4 reasoning at each decision node. GPT-4 generated binary decisions at each node to traverse the tree and classified MS diagnoses at terminal nodes. We evaluated performance against neurologist-assessed diagnoses and characterized hallucinations (non-factual, incongruent, irrelevant, over-reliant, and logical reasoning errors). Results. In this study cohort (mean age 40{+/-}13 years; 81% women) representative of the clinic population, GPT-4 performed well in predicting MS diagnosis (84% accuracy, 79% precision, 74% recall, 91% specificity) using first neurology notes. Hallucinations occurred in 32 cases (26%), most commonly incoherence (75%) and overreliance (47%). Conclusion. A structured, LLM-guided decision framework can flag MS diagnoses from early clinical documentation. Large-scale studies are needed to mitigate hallucinations, validate this approach, and test implementation in clinical settings.

8
Evaluating Goodness of Pronunciation and Phonological Posteriors as Objective Markers of Speech Severity in Motor Speech Disorders

Wang, F.; Utianski, R. L.; Duffy, J. R.; Barnard, L. R.; Botha, H.

2026-07-16 neurology 10.64898/2026.07.14.26358076 medRxiv
Top 5%
0.1%
Show abstract

This study examined the extent to which goodness of pronunciation (GoP) scores and phonological posterior probabilities capture perceptual ratings of speech severity in individuals with motor speech disorders (MSD). Speech recordings of the word catastrophe were obtained from 489 participants, including 333 neurologically typical controls and 156 individuals with MSD. GoP scores were derived using traditional acoustic features and self-supervised speech representations, including WavLM and XLS-R, across multiple modeling approaches, while phonological posterior probabilities were extracted using Phonet. Model performance was evaluated using Kendall's rank correlations, regression, and receiver operating characteristic analyses against speech-language pathologists' perceptual ratings of sound distortion and intelligibility. Both GoP and phonological posterior probabilities were significantly associated with perceptual ratings. Self-supervised speech representations substantially outperformed traditional acoustic features, with WavLM-based GoP using k-nearest neighbors achieving the strongest performance. Across correlation, regression, and classification analyses, GoP consistently outperformed phonological posterior probabilities for both sound distortion and intelligibility. Age and gender had minimal influence on model-derived measures or their relationships with perceptual ratings. These findings demonstrate the value of self-supervised GoP as an objective measure of speech impairment while highlighting the complementary role of phonological posterior probabilities in characterizing articulatory aspects of motor speech disorders.

9
Portable Ultra-Low Field MRI Deep-Learning Algorithms for White Matter Lesion Segmentation Improve Accuracy and Reflect Clinical Disability in Multiple Sclerosis

Thommana, A. A.; Donnay, C. A.; Norato, G.; Gaitan, M. I.; Griffanti, L.; Nair, G.; Reich, D. S.; Okar, S. V.

2026-07-17 neurology 10.64898/2026.07.15.26357954 medRxiv
Top 5%
0.1%
Show abstract

White matter lesion (WML) identification, assessment, and characterization using magnetic resonance imaging (MRI) are fundamental for diagnosis and monitoring of multiple sclerosis (MS). Portable ultra-low field (pULF) MRI at 64 millitesla (mT) has been shown to visualize WML with at least one dimension greater than 4 mm. An automated WML segmentation tool catered to pULF-MRI can provide standardized and accurate quantitative measurements of WML volume. In this study, we sought to investigate and compare the accuracy of machine-learning (ML) and deep-learning (DL) pULF MRI segmentation tools. Same-day paired pULF (64mT) and high-field (HF, 3T) MRI scans from 84 adults with MS or suspected-MS (mean age {+/-} SD: 48 {+/-} 13, 62 females) included T2-FLAIR and T1w images. Reference WML segmentations were manually annotated on pULF T2-FLAIR for all scans, with WML confirmed with registered HF T2-FLAIR. HF reference WML segmentations were created. Four automated segmentation methods were applied to pULF scans: Method for Inter-Modal Segmentation Analysis (MIMoSA), an ML algorithm trained on HF WML masks; WMH-SynthSeg, a convolutional neural network model with flexible segmentation capabilities across field strengths and resolution; nnU-Net, a DL algorithm trained on pULF reference WML masks; and Pseudo-Label Assisted nnU-Net (PLAn), a DL algorithm pre-trained on HF reference WML masks and refined with 64mT reference WML masks. Two models were trained with nnU-Net, one using T2-FLAIR images only (nnU-Net-FL) and one using T1w and T2-FLAIR images (nnU-Net-FL/T1). The same was done with PLAn, creating PLAn-FL and PLAn-FL/T1. The six automated WML segmentation outputs were compared to the manual segmentations to determine Dice Similarity Coefficient (DSC) scores. Associations of WML volume estimates with clinical measures were investigated. DSC scores with pULF reference WML masks from PLAn-FL (DSC mean {+/-} SD: 0.50 {+/-} 0.24) outperformed MIMoSA (0.24 {+/-} 0.20, p < 0.0001), WMH-SynthSeg (0.30 {+/-} 0.18, p < 0.0001), nnU-Net-FL (0.41 {+/-} 0.24, p < 0.0001), and nnU-Net-FL/T1 (0.41 {+/-} 0.26, p = 0.0004). Worse Expanded Disability Status Scale (EDSS) and Scripps Neurologic Rating Scale (SNRS) scores were correlated with higher WML volumes in the pULF and HF reference masks. They were also correlated with WML volumes derived from WHM-SynthSeg, nnU-Net-FL, nnU-Net-FL/T1, PLAn-FL, and PLAn-FL/T1, but not MIMoSA. After adjusting for age, WHM-SynthSeg, nnU-Net FL, nnU-Net-FL/T1, PLAn-FL, and PLAn-FL/T1 had significant associations with EDSS and SNRS scores. nnU-Net and PLAn performed best in segmenting WML on pULF-MRI at 64 mT, providing accurate quantitative estimates of WML burden. Moreover, WML volumes estimated by these algorithms were associated with clinical measures of disability, underscoring their utility for reflecting clinical and radiological disease severity. Given pULF-MRI's mobility and lower cost, these findings highlight its relevance in clinical trials, particularly in involving more participants who face logistical constraints and barriers.

10
Comparing different neuroimaging modalities for quantification of the cholinergic system in Parkinson's disease

d'Angremont, E.; Marschall, T. M.; Renken, R. J.; Sommer, I. E.

2026-07-17 neurology 10.64898/2026.07.15.26357522 medRxiv
Top 6%
0.1%
Show abstract

Introduction Parkinson's disease (PD) is a multifactorial disorder, affecting multiple neurotransmitter systems, including the cholinergic system. Cholinergic denervation is heterogeneous across patients and difficult to predict based on clinical presentation. In this study, we assessed the sensitivity of structural MRI (sMRI) and functional MRI (fMRI) to cholinergic degeneration related to PD and to cognitive functioning in PD. We compared our results to results from previously reported [18F]Fluoroethoxybenzovesamicol ([18F]FEOBV) PET imaging, which is considered the gold standard for cholinergic imaging. Methods 34 PD patients and 10 healthy controls underwent structural T1-weighted MRI. A subset of 14 patients and 9 controls also underwent resting-state fMRI. We extracted the bilateral volumes of the nucleus basalis of Meynert (NBM) from the sMRI images. Functional connectivity (FC) from the NBM to the cortex (NBM-FC) was determined using fMRI data. Principal component analysis (PCA) was applied to reduce the dimensionality of the NBM-FC images. We assessed performances for NBM-FC in distinguishing patients from controls using stepwise logistic regression. Similarly, NBM volume was used using logistic regression. Furthermore, the relation between these measures and cognitive function in several domains was investigated with (stepwise) linear regression. Leave-one-out cross validation (LOOCV) and bootstrapping was performed to assess robustness of the results. Results NBM-FC was well able to discriminate patients from controls with an AUC of 0.84 (95% CI: 0.62-1). NBM volume showed lower performance, but was still better than chance: AUC: 0.75 (95% CI: 0.57-0.93). Significant correlations were found between 1) cognition in the attentional domain and NBM-FC (r=0.63; p=.015) and 2) global cognition and NBM volume (r=0.55, p=.001). These results were inferior to those previously reported using [18F]FEOBV tracer uptake (see Chapter 6). Bootstrapping revealed that NBM volume of only the left hemisphere was stably related to PD diagnosis and global cognition in PD patients. We found that a lower NBM-FC in specific brain areas, including the fusiform gyrus, supramarginal gyrus and dorsolateral prefrontal cortex, was related to PD diagnosis. Bootstrapping revealed no stable NBM-FC pattern related to attention. Conclusion Although MRI results were slightly inferior to [18F]FEOBV PET data, MRI may provide a cheaper and more widely available alternative for cholinergic imaging. We recommend testing the utility of MRI as predictor and monitor of cholinergic treatment effect in a longitudinal study.

11
Validation of an Assessment Scale for a Low-Tech Laparoscopic Appendectomy Simulation and Its Relevance for Formative Self-Assessment

Tumameu Kouam, T. H.; Renoult, L.; Poitevin, M.; Jourdin, L.; Herve, C.; Meignan, P.; Podevin, G.; Schmitt, F.

2026-07-21 medical education 10.64898/2026.07.20.26358477 medRxiv
Top 7%
0.1%
Show abstract

Introduction: Laparoscopic appendectomy is an ideal procedure for acquiring laparoscopic skills through simulation. Nevertheless, technical training is time consuming for surgical trainers to provide constructive feedback, but this could be improved by the development of validated tools that enable appropriate formative self-assessment. For this reason, we developed a structured assessment scale for a laparoscopic appendectomy exercise using a low-fidelity simulator. The objective of this study was to validate the scale for use in formative self-assessment. Methods: During laparoscopic simulation sessions in 2025-2026, participants with varying levels of experience performed a standardized laparoscopic appendectomy (LAP) exercise on a low-fidelity simulator. Performance was assessed through formative self- and external assessment using a specific scale derived from the OSATS (Objective Structured Assessment of Technical Skills) score. Content and construct validity, internal consistency, reproducibility, and reliability in both hetero- and self-assessment were analyzed. Results: Thirty-two participants were included in the validation study of the LAP scale, including 7 medical students, 17 residents in pediatric, visceral, urological, and gynecological surgery, and 8 practicing surgeons. The content of the scale was deemed relevant by 80% of the users. It demonstrated excellent construct validity, with scores increasing according to level of experience: 9.9 +/- 0.7 among students, 12.7 +/- 3.3 among junior residents, 16.6 +/- 3.3 among experienced residents, and 18.8 +/- 0.9 among practicing surgeons (p < 0.0001). Reproducibility and internal consistency were significant, while inter and intrarater reliability were excellent (correlation coefficients r = 0.90 and 0.91; p < 0.0001), as was the correlation between external and self-assessment (r = 0.81; p < 0.0001). Self-assessment was more reliable among experienced learners than among novices. Conclusion: This standardized LAP scale is validated for both external and self-assessment, the latter requiring prior training to be reliable and formative.

12
Learned ultrasound segmentation and deformable CT fusion for augmented reality endovascular surgery

Dillon, T. M.; Quevedo Moreno, D.; Rutherford, E. K.; Ayers, B.; Salomon, B.; Kubi, B.; Thomas, J.; Roche, E.

2026-07-17 cardiovascular medicine 10.64898/2026.07.15.26358084 medRxiv
Top 7%
0.1%
Show abstract

Minimally invasive endovascular procedures offer reduced surgical trauma, shorter recovery times, and improved outcomes, but rely on 2D fluoroscopic X-ray imaging, which provides limited depth perception and exposes patients and clinicians to ionizing radiation. Here we present an augmented reality (AR) system that fuses intravascular ultrasound (IVUS) and electromagnetic (EM) position tracking with preoperative computed tomography (CT) to produce an anatomically accurate, deformation-corrected navigational reference. A robotic device performs ECG-gated pullback of the IVUS probe, capturing 4D aortic motion across the cardiac cycle. We introduce a deep learning architecture for extracting vascular lumen boundaries and side-branch orifices from artifact-prone IVUS streams, and a semantically driven non-rigid CT-IVUS fusion pipeline robust to false positive landmarks. We evaluate the platform with trained surgeons in benchtop phantom studies and in-vivo ovine models, and demonstrate its application to fenestrated endovascular aneurysm repair (FEVAR). Compared to fluoroscopy alone, AR guidance significantly reduces cannulation time, radiation exposure, and cognitive workload, while improving procedural efficiency and safety. Our IVUS-EM and CT aortic datasets are released open source.

13
Quantitative Prognostic Modeling in Aneurysmal Subarachnoid Hemorrhage: Multicenter Validation of the eSAH Score

Salman, S.; Graf von Moy, C.; Haidenberger, F.; Ahmed, M.; Foettinger, F.; Sharma, R.; Gutierrez-Aguirre, S.; de Toledo, O.; Patel, V.; Yujia-Wei, D.; Rezai Jahromi, B.; Brandmeir, N.; Lakkaraju, K.; Ombada, M.; Aguilar-Salinas, P.; Miller, D.; Erickson, B.; Hanel, R.; Tawk, R.; Byrne, R.; Freeman, W. D.

2026-07-21 neurology 10.64898/2026.07.18.26358390 medRxiv
Top 8%
0.1%
Show abstract

Background: aneurysmal subarachnoid hemorrhage (aSAH) is neurological emergency associated with substantial mortality and disability. Current grading systems such as the modified Fisher Scale (mFS) and World Federation of Neurological Societies (WFNS) score, rely on semiquantitative and examination based assessments. Hence, they demonstrate limited predictive precision. The enhanced subarachnoid hemorrhage (eSAH) score is a simplified quantitative model integrating age, Glasgow Coma Scale (GCS), and cisternal subarachnoid hemorrhage volume (SAHV) to predict clinical outcomes after aSAH. Methods: We performed a retrospective multicenter cohort study that included 1088 patients across three tertiary-care centers the United States. Predictive performance for unfavorable functional outcome, in-hospital mortality and delayed cerebral ischemia (DCI) was evaluated using receiver operating characteristic (ROC) analysis and area under the curve (AUC). Comparative analyses were performed and compared to the WFNS and mFS grading systems. Results: the eSAH score demonstrated excellent discrimination for unfavorable functional outcome at discharge ( AUC 0.89 ) and in-hospital mortality (AUC 0.87). The DCI subscore demonstrated good discriminatory performance for predicting DCI (AUC 0.77). Compared with conventional grading systems, this was superior to both the WFNS (AUC 0.75) and the mFS ( AUC 0.70). increasing eSAH scores were additionally associated with progressively higher rates of mortality and unfavorable functional outcomes. Conclusion: the eSAH score demonstrates strong external validity, reproducibility and superior predictive performance compared with conventional grading systems in a large multicenter cohort. These findings support the clinical utility of quantitative hemorrhage burden integration for early risk stratification in patients with aSAH.

14
Intravesical Lactobacillus rhamnosus GG reduces symptoms among people with spinal cord injury and disease who use intermittent catheterization: A randomized comparison of two- and four-dose regimens.

Groah, S. L.; Tractenberg, R. E.; Riegner, C. R.; Forster, C. S.

2026-07-20 urology 10.64898/2026.07.17.26358333 medRxiv
Top 8%
0.1%
Show abstract

Background: Urinary tract infection (UTI) is the most common secondary condition among people with spinal cord injury/disease (SCI/D). Intravesical Lacticaseibacillus rhamnosus GG (LGG) is an antibiotic-sparing approach to managing urinary symptoms. Objective: Determine the optimal number of doses of intravesical LGG for urinary symptom reduction. Design: Prospective, randomized, two-arm dosing trial. Setting: National recruitment with a local subsample providing urine samples in Washington, DC, USA. Participants: Adults with SCI/D and neurogenic lower urinary tract dysfunction (NLUTD) who use intermittent catheterization (IC); 177 enrolled and randomized (intention-to-treat), with 76 compliant instillers (39 low-dose, 37 high-dose) in the per-protocol analytic sample. Interventions: Two (2 doses/24 hours) or four (4 doses/36 hours) intravesical LGG regimens, self-initiated in response to cloudier or malodorous urine per the Self-Management Protocol using Probiotics (SMP-Pro). Main Outcome Measures: Primary: proportion achieving [&ge;]20% reduction on the Urinary Symptom Questionnaire for Neurogenic Bladder-Intermittent Catheter version (USQNB-IC). Secondary: urinary biomarkers (leukocyte esterase, nitrite, white blood cells, urinary neutrophil gelatinase-associated lipocalin [uNGAL]) and standard urine culture (SUC) in a local subsample. Results: By Day 2, 57.9% (63.8% low-dose; 51.2% high-dose) achieved [&ge;]20% total symptom reduction; high-dose success rose to 70.0% by Day 4. Thirty percent of high-dose participants did not respond at either time point and could not be distinguished from responders by demographics or urine biomarkers. Urinary biomarkers and SUC were unchanged pre- to post-instillation. No serious adverse events were adjudicated as attributable to intravesical LGG by an independent Data Safety Monitoring Board (DSMB). Conclusions: A two-dose course of intravesical LGG yields clinically meaningful symptom improvement in the majority of people with SCI/D and NLUTD who use IC; four doses benefits a meaningful subgroup of two-day non-responders, while a small cohort remains nonresponsive. These results provide preliminary dosing guidance and support progression to a definitive trial.

15
Diffusion MRI Profiles Map onto Distinct Inflammatory States After Adolescent Concussion: A CARE4Kids Study

Lim, A.; Gill, J. M.; Bickart, K. C.; Onicas, A. I.; Bazarian, J. K.; Alice, J.; Mac Donald, C. L.; Brown, A.; Cook, L.; Rivara, F. P.; Gioia, G. A.; Giza, C. C.; Dennis, E. L.; Concussion Assessment, Research, and Education for Kids (CARE4Kids) Consortium,

2026-07-20 neurology 10.64898/2026.07.17.26358354 medRxiv
Top 9%
0.1%
Show abstract

Importance: Neuroinflammation is a key component of the response to injury after concussion, but direct links between diffusion MRI metrics and specific plasma inflammatory pathways in human concussion have not been established. Objective: To examine associations between diffusion MRI metrics and pathway-level inflammatory proteomic signatures in adolescents during the subacute period after concussion. Design, Setting, and Participants: Cross-sectional analysis of data from the CARE4Kids Consortium, a six-site prospective study. Participants were English-speaking adolescents ages 11-17.99 with concussion and symptoms at 7-35 days post-injury. Data were collected between 2022-2024. Of 370 enrolled participants, 122 had both diffusion MRI and plasma proteomics available for analysis. Exposure: Advanced diffusion MRI metrics were converted to z-scores and participants were grouped by the spatial extent of outlier values (potholes and peaks) across 15 white matter regions of interest. Nine non-redundant groupings were selected for primary analysis. Main Outcomes and Measures: Pathway-level inflammatory profiles derived from gene set enrichment analysis (GSEA) of ~5,400 plasma proteins measured by Olink proximity extension assay, targeting nine hallmark inflammatory pathways spanning initiation through resolution. Persistent symptoms were assessed 64-115 days post-injury. Results: Diffusion metrics reflecting tissue disorganization were associated with upregulation of the coagulation pathway, consistent with hemostatic-inflammatory signaling. Metrics reflecting reduced tissue complexity and neurite density were associated with upregulation of interferon- and interferon-{gamma} response pathways, consistent with microstructural remodeling driven by cellular immune activation. Elevated free water content was associated with downregulation of most inflammatory pathways and trend-level transforming growth factor - {beta} upregulation, reflecting inflammatory resolution. Time since injury did not differ between groups based on free water (Kolmogorov-Smirnov p = 0.97), suggesting these differences reflect individual variability in recovery pace. Exploratory analyses showed a trend toward lower odds of persistent symptoms in the group with elevated free water content (odds ratio = 0.51, p = 0.18). Conclusions and Relevance: Multiple diffusion MRI metrics are differentially sensitive to distinct neuroinflammatory states in the subacute period after adolescent concussion. These findings suggest that diffusion imaging could serve as a non-invasive tool for inflammatory phenotyping, with potential implications for identifying patients who may benefit from targeted immunomodulatory intervention.

16
Toward precision rehabilitation in adolescent mild traumatic brain injury: leveraging physiologic data from commercially available smartwatches to identify patient subgroups

Kettlety, S. A.; Akrong, E. R.; Suskauer, S. J.; Roemmich, R. T.; Slomine, B. S.; Svingos, A. M.

2026-07-17 pediatrics 10.64898/2026.07.16.26358245 medRxiv
Top 9%
0.1%
Show abstract

Autonomic dysfunction is a common sequela of mild traumatic brain injury (mTBI). Physical activity progression is an integral component of mTBI rehabilitation, particularly in addressing autonomic dysfunction. However, clinicians often rely on point-in-time evaluation of orthostatic and exercise intolerance to guide activity recommendations. Commercially available wearable devices (e.g., Fitbits) provide an opportunity to evaluate heart rate response to activity in a real-world setting. Previous work has used physiologic (heart rate) and activity (step count) data to identify subgroups of adults with stroke that may be used to guide activity recommendations. This method may be useful to subgroup youth post-mTBI to identify those who have abnormal physiologic responses to activity. We aimed to identify subgroups using heart rate and step count data in adolescents presenting for specialty care after diagnosed mTBI. Eighty participants aged 13-18 within six months of mTBI diagnosis were recruited to wear a Fitbit Sense 2. Data from seven days and two nights collected within fourteen days of enrollment were included. A group-based steps per minute (SPM) threshold (25th percentile; 10 SPM) and individualized heart rate threshold (20% heart rate reserve (HRR)) were used to classify each minute of active daytime data into one of four quadrants: SPM>10 & HRR>20% (QI), SPM<10 & HRR>20% (QII), SPM<10 & HRR<20% (QIII), and SPM>10 & HRR<20% (QIV). We used percentage of minutes in each quadrant, mean steps per day, percentage of minutes with zero steps, mean SPM in QI, and resting heart rate in a k-means clustering algorithm to identify subgroups. We evaluated subgroup differences by clustering variables using Kruskal-Wallis tests. Sixty-one participants were included. Three subgroups emerged: Sedentary (n=12), Active (n=23), and Atypically Elevated Heart Rate (AEHR; n=26). Subgroups varied significantly on all clustering variables (p<0.01). The Active subgroup took a high number of steps per day, had lower sedentary time, and had the highest activity intensity (mean SPM in QI). The Sedentary subgroup took fewer steps per day compared to the Active subgroup, had high sedentary time, and showed the highest resting heart rate. The AEHR subgroup took fewer steps per day compared to the Active subgroup and had high sedentary time. The AEHR subgroup also spent a higher percentage of time with an atypically high heart rate response to low levels of activity compared to the other subgroups. Our findings suggest that data from wearable devices can identify subgroups of adolescents with mTBI with distinct physiologic/physical activity profiles, which may ultimately be used to inform personalized activity prescriptions. Future work should aim to understand how the identified subgroups relate to longitudinal outcomes.

17
Genome-Wide Association Studies and Deep-Learning Functional Annotation of Opioid Use Disorder across Three Ancestries in the All of Us Research Program

Gu, S.; Petrovitch, D.; Hall, O. T.; Lambert, J. W.; Kember, R. L.; Nahid, N. A.; Ma, Q.; Sprague, J. E.; McDonough, C. W.; Johnson, J. A.

2026-07-17 addiction medicine 10.64898/2026.07.15.26358096 medRxiv
Top 9%
0.0%
Show abstract

Background: Opioid use disorder (OUD) is heritable, yet most genome-wide association studies (GWAS) have focused on European populations, leaving the genetic architecture of OUD in non-European populations underexplored. Methods: We conducted GWAS of OUD across three ancestries using electronic health records and genomic data from 52,357 All of Us Research Program participants (8,912 cases; 43,445 matched opioid-exposed controls; 48.5% female). Participants were stratified into European (EUR), African (AFR), and Admixed American (AMR) ancestry groups for logistic regression GWAS, with independent replication in the Million Veteran Program. We then applied the deep-learning model AlphaGenome to predict the tissue-specific transcriptomic and splicing consequences of top risk variants across 13 reward-pathway brain regions. Results: We identified and replicated a novel DDX6 risk locus, alongside established OPRM1 and FURIN signals. AlphaGenome predicted the DDX6 regulatory allele downregulates the stress-resistance gene FOXR1 in the nucleus accumbens, while the protective OPRM1 variant (rs1799971) upregulates OPRM1 expression across reward networks. Other signals of interest included IL6R and SHISA9 (EUR); GHR (AFR); and ASTN2 (AMR). Conclusions: This study identifies DDX6 as a novel OUD risk locus, replicates associations with OPRM1 and FURIN, and highlights biologically plausible ancestry-specific signals in AFR and AMR populations. We also replicated top variants in an independent population. Finally, integrating GWAS with deep-learning annotations provides specific, localized biological hypotheses to guide future experimental validation and targeted therapeutics.

18
Rationale and guidance for implementing the continual reassessment method for dose-finding in controlled human infection model studies

Weerasinghe, C.; Osowicki, J.; Simpson, J. A.; Crocker-Buque, T.; McCarthy, J.; Williams, E.; Price, D. J.

2026-07-17 infectious diseases 10.64898/2026.07.16.26358128 medRxiv
Top 9%
0.0%
Show abstract

Controlled human infection models (CHIMs) are increasingly used in infectious disease research to study pathogen dynamics and evaluate interventions under controlled conditions. However, these studies are resource-intensive and involve ethical and safety constraints, making efficient study design critical. Dose-finding is a key early component in CHIMs, where the aim is to identify a challenge dose that achieves a target infection probability. Traditional rule-based designs are commonly used but can be inefficient, motivating the use of model-based adaptive approaches such as the Bayesian Continual Reassessment Method (CRM). Although CRM has been extensively studied and widely adopted in Phase I oncology trials for identifying the maximum tolerated dose of therapeutics, its application in CHIM settings remains limited, particularly when the endpoint of interest is infection. This tutorial provides step-by-step guidance for implementing a Bayesian CRM in dose-finding CHIMs, using an oropharyngeal Neisseria gonorrhoeae challenge as a motivating case study. The framework outlines key design components, including dose-grid specification, dose-response model, prior elicitation, Bayesian updating, decision rules, and stopping criteria, with particular emphasis on a clinically interpretable parameterisation. Trial operating characteristics are evaluated through simulation studies under multiple dose-response scenarios and prior-predictive analyses, and compared with a commonly used '3+3' type rule-based design. This work highlights the advantages of Bayesian model-based designs for dose-finding in CHIMs over classic rule-based designs and provides a structured, reproducible framework for implementing CRM, supporting their application in future CHIM studies.

19
Efficient stochastic epidemic simulation via the Sellke construction

van Boven, M.; Bootsma, M. C.

2026-07-17 epidemiology 10.64898/2026.07.16.26358219 medRxiv
Top 9%
0.0%
Show abstract

Stochastic epidemic models are a cornerstone of infectious disease epidemiology and are often used to study intervention scenarios. However, large run-to-run variability can make intervention effects difficult to estimate precisely. We revisit the epidemic Sellke construction, which assigns each individual an infection threshold for the cumulative infection hazard such that, conditional on the thresholds, the epidemic trajectory becomes deterministic. This enables coupling of simulations with and without an intervention, yielding low-variance effect estimates even when outcomes such as final size or peak incidence vary widely between runs. We develop an exact, event-driven implementation that maintains infection and recovery events in priority queues. Cumulative infection-hazard updates require O(log N) time per event, yielding overall complexity O(Elog N) for E events in a population of size N. The implementation achieves computational performance comparable to the classical Gillespie algorithm while naturally accommodating non-Markovian infectious periods and complex infectiousness profiles. We illustrate the approach using distance-dependent spread of avian influenza between poultry farms in the Netherlands and a multilayer population with households, schools, and workplaces. In both examples, coupling enables efficient within-run comparisons of intervention scenarios across stochastic realisations.

20
Bridging surveillance gaps in dengue: a hierarchical model integrating mixed data sources for transmission estimation and vaccine targeting

Djaafara, B. A.; Elyazar, I. R.; Yosephine, P.; Surya, A.; Silalahi, F. S.; Handito, A.; Thohir, B.; Aryani, D.; Gunawan, D.; Nisa, A. K.; Prianto, E.; Samad, I.; Cook, A. R.; Huang, A. T.; Clapham, H. E.; Bhatt, S.; Mishra, S.

2026-07-17 epidemiology 10.64898/2026.07.15.26358208 medRxiv
Top 9%
0.0%
Show abstract

Estimating dengue force of infection (FOI) is essential for understanding transmission dynamics and targeting intervention programmes, yet surveillance data in endemic settings required for estimations are often incomplete, with varying formats. We developed a Bayesian hierarchical catalytic model that jointly fits age-stratified case data, aggregate case data, and seroprevalence surveys within a single framework, incorporating external covariates to improve parameter identifiability. Synthetic validation showed that covariates alone recovered accurate FOI point estimates even when most districts contributed only aggregate data, but did so with poorly calibrated uncertainty; anchoring the model with a single seroprevalence survey was necessary to bring credible interval coverage close to nominal. Applied to 128 districts across Java and Bali, Indonesia (2016-2024), the model revealed substantial spatial heterogeneity in FOI and reporting rates. Many districts in Java exceeded the WHO-suggested seroprevalence threshold for vaccine introduction, yet were classified as low-priority when using reported incidence as prioritisation criterion, particularly in areas with weak surveillance. Model-based seroprevalence estimation, integrating multiple data sources, offers a more consistent basis for identifying high-priority districts for vaccine introduction, and is less susceptible to surveillance bias than reported incidence.